Goto

Collaborating Authors

 genetic marker


Soybean Researcher Uses Drones to Aid Genetics Analysis

#artificialintelligence

High throughput genetic analysis is a tool that allows researchers to analyze a lot of DNA data in a short period of time. The work is commonly done in a lab with scientific instruments. Larry Purcell uses it to evaluate thousands of agricultural test plots at once. He does it from a distance of 100 feet -- straight up. Using an off-the-shelf aerial drone, Purcell can identify those soybean plants that have the genetic make-up, or genotype, for high rates of nitrogen fixation.


Large-Scale Local Causal Inference of Gene Regulatory Relationships

arXiv.org Machine Learning

Gene regulatory networks play a crucial role in controlling an organism's biological processes, which is why there is significant interest in developing computational methods that are able to extract their structure from high-throughput genetic data. Many of these computational methods are designed to infer individual regulatory relationships among genes from data on gene expression. We propose a novel efficient Bayesian method for discovering local causal relationships among triplets of (normally distributed) variables. In our approach, we score covariance structures for each triplet in one go and incorporate available background knowledge in the form of priors to derive posterior probabilities over local causal structures. Our method is flexible in the sense that it allows for different types of causal structures and assumptions. We apply our approach to the task of learning causal regulatory relationships among genes. We show that the proposed algorithm produces stable and conservative posterior probability estimates over local causal structures that can be used to derive an honest ranking of the most meaningful regulatory relationships. We demonstrate the stability and efficacy of our method both on simulated data and on real-world data from an experiment on yeast. Introduction Gene regulatory networks (GRNs) play a crucial role in controlling an organism's biological processes, such as cell differentiation and metabolism [1]. If we knew the structure of a GRN, we could intervene in the developmental process of the organism, for instance by targeting a specific gene with drugs. This manuscript version is made available under the CC-BY-NC-ND 4.0 license https://creativecommons.org/licenses/by-nc-nd/4.0/ . Gene regulatory relationships are inherently causal: one can manipulate the expression level of one gene (the'cause') to regulate that of another gene (the'effect'). Because of this, many GRN inference algorithms rely on causal modeling. Causal networks such as GRNs can be inferred globally or locally.


Teen Inventor Designs Noninvasive Allergy Screen Using Genetics and Machine Learning

#artificialintelligence

One of Ayush Alag's earliest memories is of biting into a chocolate bar with cashew nuts and suddenly feeling his throat get itchy. For most of his childhood, the Santa Clara, California resident avoided eating anything with cashews and other nuts that caused irritation as best as he could. By his middle school years, he and his parents wanted to know for sure: did he have a serious food allergy, like 32 million other Americans, or was it just a food sensitivity? They sought the help of an allergist, Joseph Hernandez of Stanford University. Hernandez told them that the difference between an allergy and a food sensitivity is huge.


Crop Yield Prediction Using Deep Neural Networks

arXiv.org Machine Learning

Crop yield is a highly complex trait determined by multiple factors such as genotype, environment, and their interactions. Accurate yield prediction requires fundamental understanding of the functional relationship between yield and these interactive factors, and to reveal such relationship requires both comprehensive datasets and powerful algorithms. In the 2018 Syngenta Crop Challenge, Syngenta released several large datasets that recorded the genotype and yield performances of 2,267 maize hybrids planted in 2,247 locations between 2008 and 2016 and asked participants to predict the yield performance in 2017. As one of the winning teams, we designed a deep neural network (DNN) approach that took advantage of state-of-the-art modeling and solution techniques. Our model was found to have a superior prediction accuracy, with a root-mean-square-error (RMSE) being 12% of the average yield and 50% of the standard deviation for the validation dataset using predicted weather data. With perfect weather data, the RMSE would be reduced to 11% of the average yield and 46% of the standard deviation. Our computational results suggested that this model significantly outperformed other popular methods such as Lasso, shallow neural networks (SNN), and regression tree (RT).


Integrating omics and MRI data with kernel-based tests and CNNs to identify rare genetic markers for Alzheimer's disease

arXiv.org Machine Learning

For precision medicine and personalized treatment, we need to identify predictive markers of disease. We focus on Alzheimer's disease (AD), where magnetic resonance imaging scans provide information about the disease status. By combining imaging with genome sequencing, we aim at identifying rare genetic markers associated with quantitative traits predicted from convolutional neural networks (CNNs), which traditionally have been derived manually by experts. Kernel-based tests are a powerful tool for associating sets of genetic variants, but how to optimally model rare genetic variants is still an open research question. We propose a generalized set of kernels that incorporate prior information from various annotations and multi-omics data. In the analysis of data from the Alzheimer's Disease Neuroimaging Initiative (ADNI), we evaluate whether (i) CNNs yield precise and reliable brain traits, and (ii) the novel kernel-based tests can help to identify loci associated with AD. The results indicate that CNNs provide a fast, scalable and precise tool to derive quantitative AD traits and that new kernels integrating domain knowledge can yield higher power in association tests of very rare variants.


Predictor Variable Prioritization in Nonlinear Models: A Genetic Association Case Study

arXiv.org Machine Learning

The central aim in this paper is to address variable selection questions in nonlinear and nonparametric regression. Motivated by statistical genetics, where nonlinear interactions are of particular interest, we introduce a novel, interpretable, and computationally efficient way to summarize the relative importance of predictor variables. Methodologically, we develop the "RelATive cEntrality" (RATE) measure to prioritize candidate genetic variants that are not just marginally important, but whose associations also stem from significant covarying relationships with other variants in the data. We illustrate RATE through Bayesian Gaussian process regression, but the methodological innovations apply to other nonlinear methods. It is known that nonlinear models often exhibit greater predictive accuracy than linear models, particularly for phenotypes generated by complex genetic architectures. With detailed simulations and an Arabidopsis thaliana QTL mapping study, we show that applying RATE enables an explanation for this improved performance.


Bayesian inference. : Probabilistic machine learning and artificial intelligence : Nature : Nature Research

#artificialintelligence

A simple example of Bayesian inference applied to a medical diagnosis problem. Here the problem is diagnosing a rare disease using information from the patient's symptoms and, potentially, the patient's genetic marker measurements, which indicate predisposition (gen pred) to this disease. In this example, all variables are assumed to be binary. The relationships between variables are indicated by directed arrows and the probability of each variable given other variables they directly depend on is also shown. Yellow nodes denote measurable variables, whereas green nodes denote hidden variables.


Dogs go through a 'stroppy teenage phase'

Daily Mail - Science & tech

We all know how difficult, distracted and stroppy some teenagers can be. Now, new research has shown that dogs go through a similar'teenage' growing-up phase when they hit 8 months. The discovery comes after owners of hundreds of dogs that were tracked as they grew up reported'adolescent' behaviour at eight months. Teenagers can be difficult, distracted and stroppy. The research was aimed at spotting those that would be suitable for training as guide dogs.


Study pinpoints genetic marker that makes dogs social

Daily Mail - Science & tech

While the bond between humans and dogs now seems a natural part of life, a look at their closest living ancestor is a reminder that things weren't always that way. A new study has identified a genetic marker in dogs that sets them apart from wolves when it comes to human interaction, suggesting dogs developed a genetic condition through domestication that causes them to be so sociable. According to the researchers, this marker is the same found in people with Williams-Beuren syndrome – a condition which essentially causes people to love everyone. A new study has identified a genetic marker in dogs that sets them apart from wolves when it comes to human interaction, suggesting dogs may have developed a genetic condition through domestication that causes them to be so sociable. A genetic analysis of the world's oldest known dog remains has revealed that dogs were domesticated in a single event by humans living in Eurasia.


Cynomix Advanced Malware Analysis Technology

#artificialintelligence

Cynomix is an advanced technology developed for four years under DARPA's Cyber Genome program. It was evaluated by DARPA and MIT Lincoln Labs, and rated as the highest among all DARPA teams in its category. The goal of DARPA's Cyber Genome program was to map the genome for malware, under the premise that while over 300,000 malware strains are released daily, most are variants of a manageable number of malware families. Cynomix was conceived as a technology for identifying the unique genetic markers held in common for each malware family, and for clustering them using machine learning algorithms applied to big data sets. These algorithms cluster thousands of labeled malware ingested daily, which enables Cynomix to stay current with the newest emerging threats. This approach gives Cynomix unmatched powers of detection by analyzing a broad sampling of malware in the wild, without having to see every minor malware variation.